Papers with word error rates

18 papers
How Important is a Language Model for Low-resource ASR? (2024.findings-acl)

Copied to clipboard

Challenge: Using an n-gram language model in ASR may seem obvious, but its absence in most implementations suggests otherwise.
Approach: They examine whether using an n-gram language model in ASR can improve accuracy in low-resource languages.
Outcome: The proposed model is absent in most implementations, but it does improve accuracy in English and Mandarin.
Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech (2022.acl-long)

Copied to clipboard

Challenge: Experimental results show that the proposed model improves naturalness and prosody diversity with clear margins.
Approach: They propose a cross-utterance conditional VAE to estimate posterior probability distribution of latent prosody features for each phoneme by conditioning on acoustic features, speaker information, and text features from past and future sentences.
Outcome: The proposed model improves naturalness and prosody diversity with clear margins.
AMPS: ASR with Multimodal Paraphrase Supervision (2025.naacl-short)

Copied to clipboard

Challenge: Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition systems.
Approach: They propose a technique that augments a multilingual multimodal ASR system with paraphrase-based supervision for improved conversational ASR in multiple languages.
Outcome: The proposed technique reduces word error rates by up to 5% on a state-of-the-art multimodal model .
STORiCo: Storytelling TTS for Hindi with Character Voice Modulation (2024.eacl-short)

Copied to clipboard

Challenge: Existing datasets for read speech for Hindi lack expressiveness and character voice consistency.
Approach: They propose to use a Hindi text-to-speech (TTS) dataset to train a multi-speaker model on the single-sector data and propose to improve expressiveness and character voice consistency.
Outcome: The proposed model improves expressiveness and character voice consistency compared to the baseline single-speaker model.
Synthetic Doctor-Patient Dialogue Generation for Robust Medical ASR: A Scalable Pipeline for Vocabulary Expansion and Privacy Preservation (2026.eacl-industry)

Copied to clipboard

Challenge: Existing ASR models struggle with high word error rates (WER) on clinical vocabulary, especially medication names.
Approach: They propose to generate doctor-patient dialogues in both text and audio formats using a curated set of over 124,000 medical terms.
Outcome: The proposed pipeline generated over 1 billion audios with ground truth transcriptions.
An (unhelpful) guide to selecting the best ASR architecture for your under-resourced language (2023.acl-short)

Copied to clipboard

Challenge: English ASR now has word error rates comparable to that of human transcriptionists, but only for the handful of the world's 7000 languages with abundant training resources.
Approach: They propose to use four of the most popular ASR toolkits to train ASR models for eleven languages with limited ASR training resources: eleven widely spoken languages of Africa, Asia, and South America, one endangered language of Central America, and three critically endangered languages of North America.
Outcome: The proposed architecture outperforms four of the most popular ASR toolkits for eleven languages with limited training resources.
An Evaluation of Croatian ASR Models for Čakavian Transcription (2024.lrec-main)

Copied to clipboard

Challenge: akavian is an endangered language closely related to Croatian .
Approach: They evaluate four currently available automatic speech recognition systems that are trained on standard Croatian data and assess their performance in the transcription of akavian audio data.
Outcome: The proposed models perform better than the standard conformer model and the best-performing variant of the CTC-based model.
A Speech Recognizer for Frisian/Dutch Council Meetings (2022.lrec-1)

Copied to clipboard

Challenge: During council meetings both Frisian and Dutch are spoken, and code switching between both languages shows up frequently.
Approach: They develop a bilingual Frisian/Dutch speech recognizer for council meetings in Fryslân (the Netherlands) based on an existing Frisian and Dutch speech recognized by FAME!, which was trained and tested on radio broadcasts.
Outcome: The new recognizer is based on an existing speech recognizer for Frisian and Dutch named FAME!, which was trained and tested on radio broadcasts.
Graph-Based Phonetic Error Correction of Noisy ASR (2026.acl-industry)

Copied to clipboard

Challenge: Automatic speech recognition systems produce residual transcription errors that affect semantically critical tokens.
Approach: They propose a phonetic-based algorithm that combines phonetic graph modeling with contextual language understanding to improve automatic speech recognition.
Outcome: The proposed framework decouples phonetic reasoning from contextual semantic selection and improves accuracy.
Automatic Speech Recognition for Gascon and Languedocian Variants of Occitan (2024.lrec-main)

Copied to clipboard

Challenge: a new system for automatic speech recognition is being developed for two main Occitan dialects . the difficulty lies in the fact that Occitian is a less-resourced language .
Approach: They propose to develop an automatic speech recognition system for two Occitan dialects . they use Kaldi, acoustic models, and Whisper to create a model from corpora .
Outcome: The proposed system is based on Kaldi and Whisper for two main Occitan dialects . the system is more robust than previous systems, and the results are promising .
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models fail to adapt to unfamiliar speakers and language varieties . however, there are significant gaps in the adaptation of certain varieties based on the test speaker, variety, or recording conditions .
Approach: They propose a framework that allows for in-context learning in Phi-4 Multimodal . they find that as few as 12 example utterances reduce word error rates by 19.7% .
Outcome: The proposed framework reduces word error rates by 19.7% across diverse English corpora.
Multi-Input Attention for Unsupervised OCR Correction (P18-1)

Copied to clipboard

Challenge: Existing methods for OCR correction are mostly supervised methods that correct recognition errors in a single output.
Approach: They propose a sequence-to-sequence model with attention and a decoder with attention averaging to search for consensus among multiple sequences.
Outcome: The proposed methods cut the character and word error rates nearly in half on single inputs and can rival supervised methods.
Beyond Common Words: Enhancing ASR Cross-Lingual Proper Noun Recognition Using Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: In this work, we address the challenge of cross-lingual proper noun recognition in automatic speech recognition systems where proper nodes in an utterance may originate from a language different from the language in which the ASR system is trained.
Approach: They propose a dictionary-based method to correct ASR predictions in a large language model .
Outcome: The proposed method significantly reduces word error rates across cross-lingual proper noun recognition tasks involving three secondary languages.
Automatic Speech Recognition in Sanskrit: A New Speech Corpus and Modelling Insights (2021.findings-acl)

Copied to clipboard

Challenge: In this paper, we propose the first large scale study of automatic speech recognition in Sanskrit . we focus on the impact of unit selection in San's ASR systems .
Approach: They propose a large scale study of automatic speech recognition in Sanskrit . they propose syllable level unit selection that captures character sequences .
Outcome: The proposed model captures character sequences from one vowel in the word to the next vowela.
Multilingual Models for ASR in Chibchan Languages (2024.naacl-long)

Copied to clipboard

Challenge: Existing algorithms for low resource-intensive languages are not available for these languages . a paper comparing the performance of different models and algorithms for these extremely low resource languages is presented.
Approach: They propose to fine-tune four ASR algorithms to create monolingual models for Bribri and Cabécar . they then use the best performing algorithm to train joint and transfer learning models for both languages .
Outcome: The proposed algorithms are effective in both Bribri and Cabécar, but especially in Bribri.
Large Vocabulary Read Speech Corpora for Four Ethiopian Languages: Amharic, Tigrigna, Oromo and Wolaytta (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is one of the most important technologies to support spoken communication in modern life.
Approach: They have developed four large speech corpora for four Ethiopian languages . they have word error rates of 37.65%, 31.03%, 38.02%, 33.89% for each language .
Outcome: The proposed corpora achieve word error rates of 37.65%, 31.03%, 38.02%, 33.89% for Amharic, Tigrigna, Oromo and Wolaytta.
NeuTral Rewriter: A Rule-Based and Neural Approach to Automatic Rewriting into Gender Neutral Alternatives (2021.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen an increasing need for gender-neutral and inclusive language.
Approach: They propose a rule-based and a neural approach to gender-neutral rewriting for English . they use manually curated synthetic and natural data to train a rewriter .
Outcome: The proposed approach improves on the rule-based approach with word error rates below 0.18% on synthetic, in-domain and out-domain test sets.
Improving Speech Recognition for the Elderly: A New Corpus of Elderly Japanese Speech and Investigation of Acoustic Modeling for Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: In an aging society, a highly accurate speech recognition system is needed for use in electronic devices for the elderly but this cannot be achieved using conventional speech recognition systems due to the unique features of the speech of elderly people.
Approach: They construct a new corpus of elderly Japanese speech from existing Japanese speech corpora and train them using existing data.
Outcome: The proposed models achieve word error rates (WER) as low as 13.38%, exceeding the results of the previous study.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations